Skip to content

Run Laya encoder and decision layers on CUDA - #49

Draft
linear3735 wants to merge 28 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-encoder
Draft

linear3735 wants to merge 28 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-encoder

Conversation

@linear3735

@linear3735 linear3735 commented Sep 30, 2026 •

Copy link
Copy Markdown

Purpose

Run Laya's 28 encoder layers and two decision transformer layers from Rust. Validate the checkpoint and kernel bundle before loading, then leave hidden states on the GPU for the scorer.

Part of #14. Depends on #48 and the existing Hopper kernel build. Uses original RoPE and eager execution. Scorer, HTTP and CUDA Graph integration are separate steps.

Kept in draft until #48 merges. The diff against main currently includes its dependencies. Review this increment: 419 core/configuration lines for encoder execution.

Tests now live under the repository-root tests/ directory. Cargo target names and test coverage are unchanged.

Test Plan

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked

Compare with official Laya 0.3.20 on H800: 13 real request cases, including 16 distinct rows at 16×512. Check intermediate and final hidden states, reverse request order, and repeat each workspace without intermediate readbacks. Reproduction commands are in recipe/laya/README.md. The residency/workspace recipe covers the inherited resource checks.

System1-Omni Version / Commit: b7c9384. Encoder increment: 2f5ff4b → ff4d8fa; GPU run: 684470a. Later changes add pre-load rotary-size validation, device selection and test/documentation layout; model execution and kernels are unchanged.

Test Result

Current CPU CI: 45 tests passed; 9 checkpoint/GPU tests skipped. fmt, strict Clippy and release build passed. This documentation-only follow-up also passed local strict MkDocs and rendered-anchor checks. The historical GPU validation and benchmark below were not rerun; encoder and CUDA source are unchanged.

All 208 valid-token comparisons had zero measured error. Repeated runs produced identical output bytes. Candidate outputs were finite; empty-key padding was excluded from equality checks. These checks validate implementation parity.

CI for b7c9384: Rust CI, Docs build, benchmark harness tests passed.

H800 comparison, 2026-10-03: original RoPE in #49 vs Laya 0.3.20 eager fast path. Same GPU, checkpoint, token IDs, lengths and question types; BF16, concurrency 1, CUDA Graphs disabled. Scope: embedding + 28 encoder + 2 decision layers; scorer and HTTP are outside this benchmark.

Host p50 is 0.3–1.6% lower for short/long inputs and 0.4–0.5% higher at 16×512.

Each cell is run 1 / run 2; latencies are in ms. Order: upstream→Rust, then Rust→upstream. Each case/backend/run uses 20 warmup calls and 100 measured calls, after a feasibility pass. Host timing includes uploads and completion sync; req/s counts encoder calls. Rust includes validation/serialization and per-upload sync; upstream uses prebuilt CPU tensors and a final sync. Loading, compilation, warmup and readback are excluded.

Input Upstream p50 Upstream p95 Rust p50 Rust p95 Upstream req/s Rust req/s Rust/upstream p50
short_1 (1×48) 2.831/2.828 2.849/2.850 2.786/2.790 2.813/2.827 353.2/353.2 358.9/358.1 0.9843/0.9867
short_3 (4×64) 3.236/3.245 3.243/3.250 3.214/3.218 3.228/3.224 309.0/308.1 311.0/310.7 0.9933/0.9919
long_1 (1×464) 3.801/3.813 3.843/3.841 3.767/3.776 3.781/3.791 262.5/262.2 265.5/264.8 0.9909/0.9903
long_3 (4×464) 7.605/7.598 7.629/7.630 7.585/7.577 7.610/7.596 132.3/132.3 132.9/132.3 0.9974/0.9972
mixed_16 (16×512) 25.705/25.709 25.730/25.727 25.812/25.831 25.841/25.851 38.9/38.9 38.7/38.7 1.0042/1.0047

30/30 final-hidden comparisons passed: nRMS=0, max_abs=0; existing thresholds unchanged. These are implementation-parity checks.

CUDA-event timings and reproduction

CUDA events were measured in separate passes after uploads and before completion sync. They cover the compute-stream interval, including launch gaps. Same two rounds and units as above.

Input Upstream p50 Upstream p95 Rust p50 Rust p95
short_1 (1×48) 2.801/2.831 2.832/2.854 2.771/2.773 2.793/2.795
short_3 (4×64) 3.212/3.237 3.221/3.244 3.195/3.198 3.202/3.207
long_1 (1×464) 3.772/3.812 3.808/3.855 3.742/3.760 3.762/3.772
long_3 (4×464) 7.546/7.595 7.566/7.612 7.545/7.549 7.556/7.561
mixed_16 (16×512) 25.662/25.699 25.682/25.718 25.785/25.789 25.804/25.806

Reproduction, frozen inputs, measurement scripts and all raw timings. The two rounds remain separate; no samples were removed or pooled.

gh gist clone https://gist.github.com/linear3735/c777b0dc449676ebf1c26f0990a90b16 evidence
# Follow evidence/README.md to prepare source, checkpoint and kernel bundle.
python source/recipe/laya/native/compare_encoder.py "$LAYA_CHECKPOINT" cases.json target/release/examples/encoder_bench NEW_HOST_OUTPUT --timer host --warmup 20 --samples 100
# Switch LAYA_CUDA_LIBRARY to the event wrapper for a separate pass.
python source/recipe/laya/native/compare_encoder.py "$LAYA_CHECKPOINT" cases.json target/release/examples/encoder_bench NEW_EVENT_OUTPUT --timer cuda-event --warmup 20 --samples 100

Measurement source: ff4d8fa13b2c8d52027aa2c565a0b97940d1c0ca; the 12 Laya model/runtime files match #49 head b7c9384. Only the benchmark harness and event wrapper were added for measurement. Checkpoint: convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.

H800 UUID: GPU-9e874df2-5775-c81e-0730-d23a5f091c3d; driver 580.159.03. Python 3.10.21, Laya 0.3.20, Torch 2.11.0+cu128, TileLang 0.1.14, transformers 5.17.0, Rust 1.98.1, nvcc 13.0.88. Source, checkpoint, library and input hashes are in the evidence files.

Self-review

Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@hsliuustc0106

Copy link
Copy Markdown
Contributor

can we push this PR faster?

Copy link
Copy Markdown
Contributor

Please add a controlled H800 latency comparison against upstream Laya 0.3.20's CUDA fast path. The 208 hidden-state comparisons establish numerical parity; we also want to know whether this Rust encoder improves latency.

Keep the comparison scoped to the encoder plus decision transformer layers:

  • Compare upstream eager fast-path execution (use_graphs=False) with the Rust eager encoder on the same exact GPU, checkpoint revision, token IDs, lengths, question types and precision. Report any necessary differences. A graph-enabled upstream baseline can be a separate row.
  • Cover short and long single-question inputs, multiple questions, and the distinct-row 16 x 512 case already used for parity validation.
  • Report warm p50/p95, requests/s and the Rust/upstream latency ratio by shape. Measure the same host-call boundary on both sides, including equivalent input uploads and completion synchronization; report CUDA-event execution time separately if available. Disable intermediate capture/readbacks during timing.
  • Separate loading, compilation and warmup from measured execution. Use one feasibility run followed by two measured runs per configuration, interleave the configurations, and report variability and output parity. If the difference is inconclusive, say so.
  • Include reproducible commands, frozen SHAs, software/hardware settings and raw timing results in the PR results.

This would let us assess the performance benefit before the later scorer, HTTP and CUDA Graph integration.

@linear3735

Copy link
Copy Markdown
Author

Please add a controlled H800 latency comparison against upstream Laya 0.3.20's CUDA fast path. The 208 hidden-state comparisons establish numerical parity; we also want to know whether this Rust encoder improves latency.

Keep the comparison scoped to the encoder plus decision transformer layers:

  • Compare upstream eager fast-path execution (use_graphs=False) with the Rust eager encoder on the same exact GPU, checkpoint revision, token IDs, lengths, question types and precision. Report any necessary differences. A graph-enabled upstream baseline can be a separate row.
  • Cover short and long single-question inputs, multiple questions, and the distinct-row 16 x 512 case already used for parity validation.
  • Report warm p50/p95, requests/s and the Rust/upstream latency ratio by shape. Measure the same host-call boundary on both sides, including equivalent input uploads and completion synchronization; report CUDA-event execution time separately if available. Disable intermediate capture/readbacks during timing.
  • Separate loading, compilation and warmup from measured execution. Use one feasibility run followed by two measured runs per configuration, interleave the configurations, and report variability and output parity. If the difference is inconclusive, say so.
  • Include reproducible commands, frozen SHAs, software/hardware settings and raw timing results in the PR results.

This would let us assess the performance benefit before the later scorer, HTTP and CUDA Graph integration.

Added the H800 comparison, CUDA-event timings and reproduction steps to Test Result. Latency is close to upstream; all 30 final-hidden parity checks passed with zero error.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants